Skip to content

feat: ablation tests, JSON batch report, schema update, 98 tests - #5

Closed
cschanhniem wants to merge 4 commits into
mainfrom
feat/scoring-safety-improvements
Closed

cschanhniem wants to merge 4 commits into
mainfrom
feat/scoring-safety-improvements

Conversation

@cschanhniem

Copy link
Copy Markdown
Collaborator

Summary

  • JSON batch report: build_batch_report() now generates <report>.json alongside every markdown report, containing pipeline_version, candidate_count, selected_count, generated_at, disclaimer, score_averages, and selected_ids
  • Updated batch_report.schema.json: now requires disclaimer field and includes score_averages and selected_ids properties
  • Ablation tests (test_ablation.py): 9 tests verifying that filters are doing real work:
    • Novelty filter: near-duplicate of reference is excluded; ablating the filter causes it to be selected
    • Safety filter: stricter threshold removes more candidates; ablation includes risky sequences
    • Batch report JSON generated, validates against schema, counts match output, disclaimer present

Phase 2 ablation criteria

Test Result
Removing novelty filter includes near-duplicates ✅ Verified
Removing safety filter includes risky sequences ✅ Verified
Batch report machine-readable and schema-valid ✅ New

Test plan

  • make test — 98 tests pass
  • make demo — batch report JSON generated at outputs/demo_report.json
  • ruff check src tests — clean

…tion

- Add hydrophobic_moment() to physchem.py using Eisenberg (1984) consensus scale
  at 100°/residue helical projection; literature-cited correlate of AMP activity
- Expand activity_likeness_score() to incorporate amphipathicity (15% weight)
  with reduced charge/hydrophobicity weights to keep total at 1.0
- Add recall_at_k(), random_recall_at_k(), enrichment_factor(), benchmark_summary()
  to benchmark/evaluate.py with honest disclaimer in every output
- Add 'openamp-foundry bench baseline' CLI subcommand for pipeline vs random recall
- Add 'make bench-baseline' Makefile target
- 20 new tests: amphipathicity feature, hydrophobic moment edge cases,
  recall@k boundary conditions, enrichment factor, benchmark summary structure
- Add examples/benchmark/mixed_candidates.csv (20 sequences: 5 known-active AMPs
  + 15 non-AMP control sequences) for proper enrichment benchmarking
- Add examples/benchmark/active_labels.csv (5 known-active IDs matching above)
- Add make bench-hidden-active target using bench baseline CLI
- 12 new tests in test_hidden_active_recovery.py:
  - all positives rank in top half
  - recall@5 = 1.0 (perfect recovery)
  - enrichment factor >= 2.0 at k=5 (actual EF=4.0)
  - pipeline verdict correctly says 'outperforms random'
  - negatives score lower than positives on average
  - CLI integration test for bench baseline command
  - benchmark data integrity checks
- Pipeline achieves EF=4.0 at k=5: all 5 known AMPs recovered in top 5 of 20
  vs 25% expected from random — meets Phase 2 criterion from AGENTS.md
- Expand CI to validate evidence certificates, run leakage check, and gate
  on hidden-active EF >= 1.5 at k=5 (currently achieves 4.0)
- Add test_negative_penalization.py: 20 tests verifying that problematic
  sequences (extreme hydrophobicity, high-cysteine, purely negative charge,
  long repeat runs) score lower than known AMP-like sequences on activity,
  safety, and synthesis dimensions
- Fix all 10 ruff lint warnings (unused imports) across 7 files
- 89 tests passing, ruff clean
- Add build_batch_report() to pipeline.py — generates a machine-readable
  batch_report.json alongside the markdown report; validates against schema
- Expand batch_report.schema.json to require disclaimer, score_averages,
  and selected_ids fields
- Add test_ablation.py: 9 tests verifying that removing safety/novelty filters
  degrades selection quality (per AGENTS.md Phase 2 ablation requirement):
  - Novelty filter correctly excludes near-duplicates of references
  - Ablation of novelty filter causes near-duplicates to be selected (worse)
  - Safety filter excludes high-risk sequences; ablation includes them
  - Batch report JSON generated automatically alongside markdown
  - Batch report validates against batch_report.schema.json
  - Counts in batch report match actual output
  - Disclaimer field present and non-empty
- 98 tests passing, ruff clean
@cschanhniem

Copy link
Copy Markdown
Collaborator Author

Superseded by PR #11 (feat/integrate-all-phases), which merges all Phase 2 + Phase 3 work into a single consolidation PR with 251 tests passing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant